EDBT 2026 Demo / reviewers in the wild / expert
William R. Pearson
dblp:78/2331
· DBLP profile ↗
19ranked-venue papers
4as first author
0since 2021 · last 2016
0000-0002-0727-3680ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 3 first-authorSystems, architecture and hardware · 2Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
13 papers |
Bioinformatics and computational biology · 90% Computational science and engineering · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 100% |
Topics — the 26 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
sequence alignment |
0.4 | 6 | 2013 | Adjusting scoring matrices to correct overextended alignments · Bioinform. 2013 PSI-Search: iterative HOE-reduced profile SSEARCH searching · Bioinform. 2012 Visualization of near-optimal sequence alignments · Bioinform. 2004 |
Bioinformatics and computational biology › sequence analysis
sequence similarity search |
0.2 | 2 | 2013 | PSI-Search: iterative HOE-reduced profile SSEARCH searching · Bioinform. 2012 Adjusting scoring matrices to correct overextended alignments · Bioinform. 2013 |
Computational science and engineering › information retrieval
iterative search |
0.1 | 1 | 2012 | PSI-Search: iterative HOE-reduced profile SSEARCH searching · Bioinform. 2012 |
Bioinformatics and computational biology › protein function prediction › protein classification
protein family classification |
0.1 | 1 | 2012 | PSI-Search: iterative HOE-reduced profile SSEARCH searching · Bioinform. 2012 |
Bioinformatics and computational biology
protein sequence analysis |
0.1 | 2 | 2010 | Globally, unrelated protein sequences appear random · Bioinform. 2010 Identifying distantly related protein sequences · Comput. Appl. Biosci. 1997 |
Bioinformatics and computational biology › biological database
protein domain database |
0.1 | 1 | 2010 | RefProtDom: a protein database with improved domain boundaries and homology relationships · Bioinform. 2010 |
Bioinformatics and computational biology › sequence alignment
sequence alignment visualization |
0.0 | 1 | 2004 | Visualization of near-optimal sequence alignments · Bioinform. 2004 |
Bioinformatics and computational biology
sequence analysis |
0.0 | 2 | 2002 | Empirical determination of effective gap penalties for sequence comparison · Bioinform. 2002 Identifying distantly related protein sequences · Comput. Appl. Biosci. 1997 |
Bioinformatics and computational biology › protein sequence analysis
protein similarity search |
0.0 | 1 | 2002 | Empirical determination of effective gap penalties for sequence comparison · Bioinform. 2002 |
Information retrieval › pattern matching
sequence matching |
0.0 | 1 | 2010 | RefProtDom: a protein database with improved domain boundaries and homology relationships · Bioinform. 2010 |
Algorithms and data structures › computational biology
sequence analysis |
0.0 | 1 | 2010 | Globally, unrelated protein sequences appear random · Bioinform. 2010 |
Bioinformatics and computational biology › sequence analysis
sequence comparison |
0.0 | 4 | 1997 | No Pain and Gain - Experiences with Mentat on a Biological Application · HPDC 1992 A platform for biological sequence comparison on parallel computers · Comput. Appl. Biosci. 1991 Identifying distantly related protein sequences · Comput. Appl. Biosci. 1997 |
Bioinformatics and computational biology › sequence alignment
DNA-protein alignment |
0.0 | 1 | 1997 | Aligning a DNA sequence with a protein sequence · RECOMB 1997 |
Bioinformatics and computational biology › sequence analysis
homology detection |
0.0 | 1 | 1997 | Identifying distantly related protein sequences · Comput. Appl. Biosci. 1997 |
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection |
0.0 | 1 | 1997 | Identifying distantly related protein sequences · Comput. Appl. Biosci. 1997 |
Parallel and multicore computing
parallel computing |
0.0 | 2 | 1992 | Comparing machine-independent versus machine-specific parallelization of a software platform for biological sequence comparison · Comput. Appl. Biosci. 1992 A platform for biological sequence comparison on parallel computers · Comput. Appl. Biosci. 1991 |
Visualization and visual analytics
interactive visualization |
0.0 | 1 | 2004 | Visualization of near-optimal sequence alignments · Bioinform. 2004 |
Bioinformatics and computational biology
polymerase chain reaction |
0.0 | 1 | 1995 | A New Approach to Primer Selection in Polymerase Chain Reaction Experiments · ISMB 1995 |
Bioinformatics and computational biology › genomics
primer design |
0.0 | 1 | 1995 | A New Approach to Primer Selection in Polymerase Chain Reaction Experiments · ISMB 1995 |
Bioinformatics and computational biology › sequence alignment
local alignment |
0.0 | 1 | 1992 | Aligning two sequences within a specified diagonal band · Comput. Appl. Biosci. 1992 |
Parallel and multicore computing › parallel programming models › portable programming models
machine independent parallel programming |
0.0 | 1 | 1992 | Comparing machine-independent versus machine-specific parallelization of a software platform for biological sequence comparison · Comput. Appl. Biosci. 1992 |
Parallel and multicore computing › parallel programming models
object-oriented parallel programming |
0.0 | 1 | 1992 | No Pain and Gain - Experiences with Mentat on a Biological Application · HPDC 1992 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 1992 | No Pain and Gain - Experiences with Mentat on a Biological Application · HPDC 1992 |
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search |
0.0 | 1 | 1991 | A platform for biological sequence comparison on parallel computers · Comput. Appl. Biosci. 1991 |
Parallel and multicore computing › parallel computing › parallel bioinformatics
parallel sequence search |
0.0 | 1 | 1991 | A platform for biological sequence comparison on parallel computers · Comput. Appl. Biosci. 1991 |
Parallel and multicore computing › multiprocessor system
distributed memory parallel processors |
0.0 | 1 | 1992 | No Pain and Gain - Experiences with Mentat on a Biological Application · HPDC 1992 |
Methods — techniques the papers use, named apart from their topics
local and semi-global searches · 0.2false discovery rate analysis · 0.2SCOP · 0.2PSI-BLAST · 0.2CATH · 0.2FASTA · 0.2SSEARCH · 0.2BLAST · 0.2smith-waterman local alignment · 0.1PSI-BLAST position-specific score matrix · 0.1random sequence models · 0.1dynamic programming · 0.0linda · 0.0smith-waterman · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Applying, Evaluating and Refining Bioinformatics Core Competencies (An Update from the Curriculum Task Force of ISCB's Education Committee)abstractThe Curriculum Task Force (CTF) of ISCB’s Education Committee seeks to define curricular guidelines for those who educate or train bioinformatics professionals at all career stages. A recent report of the CTF [1] presented a draft set of bioinformatics core competencies, derived from the results of surveys of (1) core facility directors, (2) career opportunities, and (3) existing curricula.
Since the publication of its 2014 report, the CTF has focused on the application of the guidelines in varied contexts to identify areas where refinement is needed. As a first step, the task force held an open meeting at the ISMB conference in July 2014. The ideas discussed at the meeting spawned four working groups (WGs), which focus on (i) defining core competencies for specific types and levels of bioinformatics training, (ii) mapping the curriculum guidelines and competencies to existing materials in order to identify the need for development of new materials, and (iii) identifying where revision of the guidelines may be valuable. The CTF is engaging the ISCB community through open WG meetings at ISCB’s official conferences. Thus far, the WGs have convened at the ISCB Great Lakes Bioinformatics Conference (Purdue University, May 2015) and at the ISMB/ECCB Conference (Dublin, Ireland, July 2015). Additionally, the CTF held a workshop at the Annual General Meeting of the Global Organization of Bioinformatics Learning, Education and Training (Cape Town, South Africa, November 2015). Specifically, the draft competencies have been employed in a wide range of activities and contexts (see Table 1 and [2–11]), including the development of new curricula, the analysis of existing curricula, and the creation of new roles involving bioinformatics. These activities have resulted in the identification of several areas where refinement would be useful:
Table 1
Summary of the activities of the ISCB Curriculum Task Force.
Identify different levels or phases of competency. It would be helpful to define different phases of competency development, or different levels of competency appropriate for distinct roles.
Define competency profiles for disciplines that don’t fit into our current silos. Bioengineering provides an illustrative example of a discipline that requires core competency in bioinformatics but does not fit into our current categories. There are almost certainly others. It would be helpful if we could provide some guidance on how to produce ‘hybrid’ competency profiles, perhaps borrowing some competencies from the TF’s core set and others from different disciplines. The LifeTrain initiative (www.lifetrain.eu) [2, 3] is collecting competency profiles for a range of disciplines of relevance to the biomedical sciences and may provide a useful resource kit for this.
Broaden the scope of the competency profiles in response to cutting-edge and emerging research. Current areas requiring improvement include incorporating competencies that capture a fundamental understanding of the biological principles central to analyzing biomolecular data, and broadening the user WG to include applications beyond medicine.
Provide guidance on the evidence required to assess whether someone has acquired each competency. For undergraduate, Master’s and PhD programs, learning outcomes for each competency, perhaps with examples of appropriate means of assessment, would be valuable. For established professionals who need to assimilate competencies into their working lives, a different approach may be required (such as keeping a portfolio to capture evidence of competency); the CTF should seek guidance from relevant professional bodies, especially in regulated professions such as healthcare.
Provide indicative course content or examples of programs that map to the competency requirements. We do not wish to prescribe what course providers should teach or how they should teach it; however, if a course provider is designing a course to meet a specific competency requirement, it may be helpful to find examples of other programs that do this successfully. One way of achieving this is by mapping existing training content to the TF’s competencies. Another way might be to provide an indication, perhaps based on several courses, of the course content that would meet the competency requirements. This would give course providers the freedom to build their own course syllabi without having to reinvent the wheel. Initiatives to collect examples of Creative Commons (or otherwise reusable) course materials will provide an extremely valuable bank of training materials that could be mapped to the core competencies. Lonnie R. Welch, Catherine Brooksbank, Russell Schwartz, Sarah L. Morgan, Bruno A. Gaëta, Alastair M. Kilpatrick, Daniel Mietchen, Benjamin L. Moore, Nicola J. Mulder, Mark A. Pauley, William R. Pearson, Predrag Radivojac, Naomi Rosenberg, Anne G. Rosenwald, Gabriella Rustici, Tandy J. Warnow |
PLoS Comput. Biol. | 11 |
| 2013 | Adjusting scoring matrices to correct overextended alignmentsabstractMOTIVATION: Sequence similarity searches performed with BLAST, SSEARCH and FASTA achieve high sensitivity by using scoring matrices (e.g. BLOSUM62) that target low identity (<33%) alignments. Although such scoring matrices can effectively identify distant homologs, they can also produce local alignments that extend beyond the homologous regions. RESULTS: We measured local alignment start/stop boundary accuracy using a set of queries where the correct alignment boundaries were known, and found that 7% of BLASTP and 8% of SSEARCH alignment boundaries were overextended. Overextended alignments include non-homologous sequences; they occur most frequently between sequences that are more closely related (>33% identity). Adjusting the scoring matrix to reflect the identity of the homologous sequence can correct higher identity overextended alignment boundaries. In addition, the scoring matrix that produced a correct alignment could be reliably predicted based on the sequence identity seen in the original BLOSUM62 alignment. Realigning with the predicted scoring matrix corrected 37% of all overextended alignments, resulting in more correct alignments than using BLOSUM62 alone. Lauren J. Mills, William R. Pearson |
Bioinform. | 2 |
| 2012 | PSI-Search: iterative HOE-reduced profile SSEARCH searchingabstractUNLABELLED: Iterative similarity searches with PSI-BLAST position-specific score matrices (PSSMs) find many more homologs than single searches, but PSSMs can be contaminated when homologous alignments are extended into unrelated protein domains-homologous over-extension (HOE). PSI-Search combines an optimal Smith-Waterman local alignment sequence search, using SSEARCH, with the PSI-BLAST profile construction strategy. An optional sequence boundary-masking procedure, which prevents alignments from being extended after they are initially included, can reduce HOE errors in the PSSM profile. Preventing HOE improves selectivity for both PSI-BLAST and PSI-Search, but PSI-Search has ~4-fold better selectivity than PSI-BLAST and similar sensitivity at 50% and 60% family coverage. PSI-Search is also produces 2- for 4-fold fewer false-positives than JackHMMER, but is ~5% less sensitive. AVAILABILITY AND IMPLEMENTATION: PSI-Search is available from the authors as a standalone implementation written in Perl for Linux-compatible platforms. It is also available through a web interface (www.ebi.ac.uk/Tools/sss/psisearch) and SOAP and REST Web Services (www.ebi.ac.uk/Tools/webservices). Weizhong Li 0001, Hamish McWilliam, Mickael Goujon, Andrew Peter Cowley, Rodrigo Lopez, William R. Pearson |
Bioinform. | 6 |
| 2010 | RefProtDom: a protein database with improved domain boundaries and homology relationshipsabstractUNLABELLED: RefProtDom provides a set of divergent query domains, originally selected from Pfam, and full-length proteins containing their homologous domains, with diverse architectures, for evaluating pair-wise and iterative sequence similarity searches. Pfam homology and domain boundary annotations in the target library were supplemented using local and semi-global searches, PSI-BLAST searches, and SCOP and CATH classifications. AVAILABILITY: RefProtDom is available from http://faculty.virginia.edu/wrpearson/fasta/PUBS/gonzalez09a. Mileidy W. Gonzalez, William R. Pearson |
Bioinform. | 2 |
| 2010 | Globally, unrelated protein sequences appear randomabstractMOTIVATION: To test whether protein folding constraints and secondary structure sequence preferences significantly reduce the space of amino acid words in proteins, we compared the frequencies of four- and five-amino acid word clumps (independent words) in proteins to the frequencies predicted by four random sequence models. RESULTS: While the human proteome has many overrepresented word clumps, these words come from large protein families with biased compositions (e.g. Zn-fingers). In contrast, in a non-redundant sample of Pfam-AB, only 1% of four-amino acid word clumps (4.7% of 5mer words) are 2-fold overrepresented compared with our simplest random model [MC(0)], and 0.1% (4mers) to 0.5% (5mers) are 2-fold overrepresented compared with a window-shuffled random model. Using a false discovery rate q-value analysis, the number of exceptional four- or five-letter words in real proteins is similar to the number found when comparing words from one random model to another. Consensus overrepresented words are not enriched in conserved regions of proteins, but four-letter words are enriched 1.18- to 1.56-fold in alpha-helical secondary structures (but not beta-strands). Five-residue consensus exceptional words are enriched for alpha-helix 1.43- to 1.61-fold. Protein word preferences in regular secondary structure do not appear to significantly restrict the use of sequence words in unrelated proteins, although the consensus exceptional words have a secondary structure bias for alpha-helix. Globally, words in protein sequences appear to be under very few constraints; for the most part, they appear to be random. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniel T. Lavelle, William R. Pearson |
Bioinform. | 2 |
| 2010 | Improving pairwise sequence alignment accuracy using near-optimal protein sequence alignmentsabstractBACKGROUND: While the pairwise alignments produced by sequence similarity searches are a powerful tool for identifying homologous proteins - proteins that share a common ancestor and a similar structure; pairwise sequence alignments often fail to represent accurately the structural alignments inferred from three-dimensional coordinates. Since sequence alignment algorithms produce optimal alignments, the best structural alignments must reflect suboptimal sequence alignment scores. Thus, we have examined a range of suboptimal sequence alignments and a range of scoring parameters to understand better which sequence alignments are likely to be more structurally accurate. RESULTS: We compared near-optimal protein sequence alignments produced by the Zuker algorithm and a set of probabilistic alignments produced by the probA program with structural alignments produced by four different structure alignment algorithms. There is significant overlap between the solution spaces of structural alignments and both the near-optimal sequence alignments produced by commonly used scoring parameters for sequences that share significant sequence similarity (E-values < 10-5) and the ensemble of probA alignments. We constructed a logistic regression model incorporating three input variables derived from sets of near-optimal alignments: robustness, edge frequency, and maximum bits-per-position. A ROC analysis shows that this model more accurately classifies amino acid pairs (edges in the alignment path graph) according to the likelihood of appearance in structural alignments than the robustness score alone. We investigated various trimming protocols for removing incorrect edges from the optimal sequence alignment; the most effective protocol is to remove matches from the semi-global optimal alignment that are outside the boundaries of the local alignment, although trimming according to the model-generated probabilities achieves a similar level of improvement. The model can also be used to generate novel alignments by using the probabilities in lieu of a scoring matrix. These alignments are typically better than the optimal sequence alignment, and include novel correct structural edges. We find that the probA alignments sample a larger variety of alignments than the Zuker set, which more frequently results in alignments that are closer to the structural alignments, but that using the probA alignments as input to the regression model does not increase performance. CONCLUSIONS: The pool of suboptimal pairwise protein sequence alignments substantially overlaps structure-based alignments for pairs with statistically significant similarity, and a regression model based on information contained in this alignment pool improves the accuracy of pairwise alignments with respect to structure-based alignments. Michael L. Sierk, Michael E. Smoot, Ellen J. Bass, William R. Pearson |
BMC Bioinform. | 4 |
| 2004 | Visualization of near-optimal sequence alignmentsabstractMOTIVATION: Mathematically optimal alignments do not always properly align active site residues or well-recognized structural elements. Most near-optimal sequence alignment algorithms display alternative alignment paths, rather than the conventional residue-by-residue pairwise alignment. Typically, these methods do not provide mechanisms for finding effectively the most biologically meaningful alignment in the potentially large set of options. RESULTS: We have developed Web-based software that displays near optimal or alternative alignments of two protein or DNA sequences as a continuous moving picture. A WWW interface to a C++ program generates near optimal alignments, which are sent to a Java Applet, which displays them in a series of alignment frames. The Applet aligns residues so that consistently aligned regions remain at a fixed position on the display, while variable regions move. The display can be stopped to examine alignment details. Michael E. Smoot, Stephanie A. Guerlain, William R. Pearson |
Bioinform. | 3 |
| 2002 | Empirical determination of effective gap penalties for sequence comparisonabstractMOTIVATION: No general theory guides the selection of gap penalties for local sequence alignment. We empirically determined the most effective gap penalties for protein sequence similarity searches with substitution matrices over a range of target evolutionary distances from 20 to 200 Point Accepted Mutations (PAMs). RESULTS: We embedded real and simulated homologs of protein sequences into a database and searched the database to determine the gap penalties that produced the best statistical significance for the distant homologs. The most effective penalty for the first residue in a gap (q+r) changes as a function of evolutionary distance, while the gap extension penalty for additional residues (r) does not. For these data, the optimal gap penalties for a given matrix scaled in 1/3 bit units (e.g. BLOSUM50, PAM200) are q=25-0.1 * (target PAM distance), r=5. Our results provide an empirical basis for selection of gap penalties and demonstrate how optimal gap penalties behave as a function of the target evolutionary distance of the substitution matrix. These gap penalties can improve expectation values by at least one order of magnitude when searching with short sequences, and improve the alignment of proteins containing short sequences repeated in tandem. Justin T. Reese, William R. Pearson |
Bioinform. | 2 |
| 2001 | Training for bioinformatics and computational biologyabstractWilliam R. Pearson; Training for bioinformatics and computational biology , Bioinformatics, Volume 17, Issue 9, 1 September 2001, Pages 761–762, https://doi.org William R. Pearson |
Bioinform. | 1 |
| 1997 | Aligning a DNA sequence with a protein sequenceabstractWe develop several algorithms for the problem of aligning a DNA sequence with a protein sequence.Our methods account for frameshift errors, but not for introns in the DNA sequence.Thus, they are particularly appropriate for comparing a cDNA sequence that contains sequencing errors with an amino acid sequence or a protein sequence database.We describe techniques for efficient implementation, verify sufficient conditions for equivalenceof several definitions of alignment, and discuss experience with these ideas in a new release of the fasta suite of database-searching programs. Zheng Zhang 0004, William R. Pearson, Webb Miller |
RECOMB | 2 |
| 1997 | Identifying distantly related protein sequencesabstractAbstract New methods for identifying distantly related proteins can be used to confirm sequence homology when only weak sequence similarity remains. These methods improve the selectivity of sequence comparison either by calculating the statistical significance of the most similar region, or by using consensus patterns rather than simple pairwise similarity scores. William R. Pearson |
Comput. Appl. Biosci. | 1 |
| 1996 | On the Primer Selection Problem in Polymerase Chain Reaction Experiments
William R. Pearson, Gabriel Robins, Dallas E. Wrege, Tongtong Zhang |
Discret. Appl. Math. | 1 |
| 1995 | A New Approach to Primer Selection in Polymerase Chain Reaction Experiments
William R. Pearson, Gabriel Robins, Dallas E. Wrege, Tongtong Zhang |
ISMB | 1 |
| 1994 | White Paper: Designing Medical Informatics Research and Library-Resource Projects to Increase What Is LearnedabstractCareful study of medical informatics research and library-resource projects is necessary to increase the productivity of the research and development enterprise. Medical informatics research projects can present unique problems with respect to evaluation. It is not always possible to adapt directly the evaluation methods that are commonly employed in the natural and social sciences. Problems in evaluating medical informatics projects may be overcome by formulating system development work in terms of a testable hypothesis; subdividing complex projects into modules, each of which can be developed, tested and evaluated rigorously; and utilizing qualitative studies in situations where more definitive quantitative studies are impractical. William W. Stead, Robert Brian Haynes, Sherrilynne S. Fuller, Charles P. Friedman, Larry E. Travis, J. Robert Beck, Carol H. Fenichel, B. Chandrasekaran 0001, Bruce G. Buchanan, Enrique E. Abola, MaryEllen C. Sievert, Reed M. Gardner, Judith Messerle, Conrade C. Jaffe, William R. Pearson, Robert M. Abarbanel |
J. Am. Medical Informatics Assoc. | 15 |
| 1993 | No pain and gain! - experiences with Mentat on a biological applicationabstractAbstract Throughout much of the parallel processing community there is the sense that writing software for distributed‐memory parallel processors Is subject to a ‘no pain—no gain’ rule: that In order to reap the benefits of parallel computation one must first suffer the pain of converting the application to run on a parallel machine. We believe this Is the result of Inadequate programming tools and not a problem Inherent to parallel processing. We will show that one can parallelize real scientific applications and obtain good performance with little effort If the right tools are used. Our vehicle for this demonstration is a 6000‐line DNA and protein sequence comparison application that we have implemented in Mental, an object‐oriented parallel processing system for both parallel and distributed architectures. We briefly describe the application and present performance information for both the Mentat version and a hand‐coded parallel version of the application. Andrew S. Grimshaw, Emily A. West, William R. Pearson |
Concurr. Pract. Exp. | 3 |
| 1992 | No Pain and Gain - Experiences with Mentat on a Biological ApplicationabstractThroughout much of the parallel processing community there is the sense that writing software for distributed memory parallel processors is subject to a 'no pain-no gain' rule: that in order to reap the benefits of parallel computation one must first suffer the pain of converting the application to run on a parallel machine. The authors believe this is the result of inadequate programming tools and not a problem inherent to parallel processing. They show that one can parallelize real scientific applications and obtain good performance with little effort if the right tools are used. Their vehicle for this demonstration is a six-thousand line DNA and protein sequence comparison application that they have implemented in Mentat, an object-oriented parallel processing system for both parallel and distributed architectures. They briefly describe the application and present performance information for both the Mentat version and a hand-coded parallel version of the application.> Andrew S. Grimshaw, Emily A. West, William R. Pearson |
HPDC | 3 |
| 1992 | Aligning two sequences within a specified diagonal bandabstractWe describe an algorithm for aligning two sequences within a diagonal band that requires only O(NW) computation time and O(N) space, where N is the length of the shorter of the two sequences and W is the width of the band. The basic algorithm can be used to calculate either local or global alignment scores. Local alignments are produced by finding the beginning and end of a best local alignment in the band, and then applying the global alignment algorithm between those points. This algorithm has been incorporated into the FASTA program package, where it has decreased the amount of memory required to calculate local alignments from O(NW) to O(N) and decreased the time required to calculate optimized scores for every sequence in a protein sequence database by 40%. On computers with limited memory, such as the IBM-PC, this improvement both allows longer sequences to be aligned and allows optimization within wider bands, which can include longer gaps. Kun-Mao Chao, William R. Pearson, Webb Miller |
Comput. Appl. Biosci. | 2 |
| 1992 | Comparing machine-independent versus machine-specific parallelization of a software platform for biological sequence comparisonabstractA platform program that performs biological sequence comparison provides a case study to compare the relative advantages of a machine-independent approach to parallel computation versus a machine-specific approach. The program consists of two routines: (i) PSCANLIB, which compares a single biological sequence against a database of sequences, and (ii) PCOMPLIB, which compares a database of sequences against another database of sequences, or against itself. The program was first parallelized to run on the Intel Hypercube parallel computer using native Hypercube commands to coordinate the parallel computation. The parallelization logic of the program was then translated into a machine-independent parallel programming language, Linda. These two approaches to parallelization are contrasted in terms of: (i) the expressive power of the logic that coordinates the parallel computation, (ii) the portability of the machine-independent version to other parallel machines and (iii) the relative efficiency of the two versions of the program. In the benchmark tests reported, the benefits of the machine-independent approach were achieved with only a modest sacrifice in efficiency. Perry L. Miller, Prakash M. Nadkarni, William R. Pearson |
Comput. Appl. Biosci. | 3 |
| 1991 | A platform for biological sequence comparison on parallel computersabstractWe have written two programs for searching biological sequence databases that run on Intel hypercube computers. PSCANLIB compares a single sequence against a sequence library, and PCOMPLIB compares all the entries in one sequence library against a second library. The programs provide a general framework for similarity searching; they include functions for reading in query sequences, search parameters and library entries, and reporting the results of a search. We have isolated the code for the specific function that calculates the similarity score between the query and library sequence; alternative searching algorithms can be implemented by editing two files. We have implemented the rapid FASTA sequence comparison algorithm and the more rigorous Smith-Waterman algorithm within this framework. The PSCANLIB program on a 16 node iPSC/2 80386-based hypercube can compare a 229 amino acid protein sequence with a 3.4 million residue sequence library in approximately 16 s with the FASTA algorithm. Using the Smith-Waterman algorithm, the same search takes 35 min. The PCOMPLIB program can compare a 0.8 million amino acid protein sequence library with itself in 5.3 min with FASTA on a third-generation 32 node Intel iPSC/860 hypercube. A. S. Deshpande, Dana S. Richards, William R. Pearson |
Comput. Appl. Biosci. | 3 |