VLDB 2026 Research / reviewers in the wild / expert
Peter Willett 0002
dblp:w/PeterWillett
· DBLP profile ↗
41ranked-venue papers
6as first author
0since 2021 · last 2019
0000-0003-4591-7173ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 27 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 9Systems, architecture and hardware · 3Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% | |
| Databases, data mining, and information retrieval
7 papers |
Information retrieval · 90% Data mining · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 60% Parallel and multicore computing · 40% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › molecular informatics
cheminformatics |
0.0 | 1 | 2004 | Chemoinformatics: an application domain for information retrieval techniques · SIGIR 2004 |
Bioinformatics and computational biology › sequence analysis
similarity search |
0.0 | 1 | 2004 | Chemoinformatics: an application domain for information retrieval techniques · SIGIR 2004 |
Bioinformatics and computational biology
biomedical text mining |
0.0 | 1 | 2003 | Protein Structures and Information Extraction from Biological Texts: The PASTA System · Bioinform. 2003 |
Bioinformatics and computational biology
drug discovery |
0.0 | 1 | 1999 | Computational analysis of molecular diversity for drug discovery · RECOMB 1999 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecular diversity |
0.0 | 1 | 1999 | Computational analysis of molecular diversity for drug discovery · RECOMB 1999 |
Information retrieval
evaluation |
0.0 | 3 | 1994 | On the Measurement of Inter-Linker Consistency and Retrieval Effectiveness in Hypertext Databases · SIGIR 1994 Searching for Historical Word-Forms in a Database of 17th-Century English Text Using Spelling-Correction Methods · SIGIR 1992 Criteria for the Selection of Search Strategies in Best-Match Document-Retrieval Systems · Int. J. Man Mach. Stud. 1986 |
Information retrieval › retrieval models
ranked retrieval |
0.0 | 1 | 2004 | Chemoinformatics: an application domain for information retrieval techniques · SIGIR 2004 |
Information retrieval
relevance feedback |
0.0 | 1 | 2004 | Chemoinformatics: an application domain for information retrieval techniques · SIGIR 2004 |
Information retrieval › retrieval models
graph-based retrieval |
0.0 | 1 | 1994 | On the Measurement of Inter-Linker Consistency and Retrieval Effectiveness in Hypertext Databases · SIGIR 1994 |
Information retrieval › document retrieval › structured document retrieval
hypertext retrieval |
0.0 | 1 | 1994 | On the Measurement of Inter-Linker Consistency and Retrieval Effectiveness in Hypertext Databases · SIGIR 1994 |
Information retrieval › evaluation
retrieval effectiveness |
0.0 | 1 | 1994 | On the Measurement of Inter-Linker Consistency and Retrieval Effectiveness in Hypertext Databases · SIGIR 1994 |
Information retrieval › document retrieval › domain-specific retrieval
historical text retrieval |
0.0 | 1 | 1992 | Searching for Historical Word-Forms in a Database of 17th-Century English Text Using Spelling-Correction Methods · SIGIR 1992 |
Information retrieval › query understanding
spelling correction |
0.0 | 1 | 1992 | Searching for Historical Word-Forms in a Database of 17th-Century English Text Using Spelling-Correction Methods · SIGIR 1992 |
Data mining › clustering
document clustering |
0.0 | 2 | 1987 | Non-Hierarchic Document Clustering Using the ICL Distributed Array Processor · SIGIR 1987 Hierarchic Document Clustering Using Ward's Method · SIGIR 1986 |
Parallel and multicore computing
parallel algorithms |
0.0 | 1 | 1987 | Non-Hierarchic Document Clustering Using the ICL Distributed Array Processor · SIGIR 1987 |
Data mining › clustering
hierarchical clustering |
0.0 | 1 | 1986 | Hierarchic Document Clustering Using Ward's Method · SIGIR 1986 |
Information retrieval
retrieval models |
0.0 | 1 | 1986 | Criteria for the Selection of Search Strategies in Best-Match Document-Retrieval Systems · Int. J. Man Mach. Stud. 1986 |
Information retrieval › interactive information retrieval › search interaction
search strategy |
0.0 | 1 | 1986 | Criteria for the Selection of Search Strategies in Best-Match Document-Retrieval Systems · Int. J. Man Mach. Stud. 1986 |
Storage systems
data compression |
0.0 | 1 | 1986 | Compression of nucleic acid and protein sequence data · Comput. Appl. Biosci. 1986 |
Storage systems › data compression
run-length coding |
0.0 | 1 | 1986 | Compression of nucleic acid and protein sequence data · Comput. Appl. Biosci. 1986 |
Storage systems › file systems
file organization |
0.0 | 1 | 1990 | Parallel Text Searching in Serial Files Using a Processor Farm · SIGIR 1990 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 1990 | Parallel Text Searching in Serial Files Using a Processor Farm · SIGIR 1990 |
Information retrieval › similarity search
nearest neighbor search |
0.0 | 1 | 1986 | Hierarchic Document Clustering Using Ward's Method · SIGIR 1986 |
Methods — techniques the papers use, named apart from their topics
named entity recognition · 0.0information extraction · 0.0genetic algorithm · 0.0hypertext link analysis · 0.0non-phonetic coding · 0.0n-gram matching · 0.0dynamic programming · 0.0single-pass clustering · 0.0run-length coding · 0.0reallocation clustering · 0.0n-gram coding · 0.0ward's method · 0.0nearest neighbor search · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Motivations, understandings, and experiences of open-access mega-journal authors: Results of a large-scale surveyabstractOpen‐access mega‐journals (OAMJs) are characterized by their large scale, wide scope, open‐access (OA) business model, and “soundness‐only” peer review. The last of these controversially discounts the novelty, significance, and relevance of submitted articles and assesses only their “soundness.” This article reports the results of an international survey of authors (n = 11,883), comparing the responses of OAMJ authors with those of other OA and subscription journals, and drawing comparisons between different OAMJs. Strikingly, OAMJ authors showed a low understanding of soundness‐only peer review: two‐thirds believed OAMJs took into account novelty, significance, and relevance, although there were marked geographical variations. Author satisfaction with OAMJs, however, was high, with more than 80% of OAMJ authors saying they would publish again in the same journal, although there were variations by title, and levels were slightly lower than subscription journals (over 90%). Their reasons for choosing to publish in OAMJs included a wide variety of factors, not significantly different from reasons given by authors of other journals, with the most important including the quality of the journal and quality of peer review. About half of OAMJ articles had been submitted elsewhere before submission to the OAMJ with some evidence of a “cascade” of articles between journals from the same publisher. Simon Wakeling, Claire Creaser, Stephen Pinfield, Jenny Fry, Valérie Spezi, Peter Willett 0002, Monica Lestari Paramita |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2013 | Drug screening with Elastic-net multiple kernel learningabstractWe apply Elastic-net Multiple Kernel Learning (MKL) to the MDL Drug Data Report (MDDR) database for the problem of drug screening. We show that combining a set of kernels constructed from fingerprint descriptors, can significantly improve the accuracy of prediction, against a Support Vector Machine trained on each kernel separately. To the best of our knowledge, this is the first application of MKL to the MDDR database for drug screening. Kitsuchart Pasupa, Zakria Hussain, John Shawe-Taylor, Peter Willett 0002 |
BIBE | 4 |
| 2011 | Novel base triples in RNA structures revealed by graph theoretical searching methodsabstractBACKGROUND: Highly hydrogen bonded base interactions play a major part in stabilizing the tertiary structures of complex RNA molecules, such as transfer-RNAs, ribozymes and ribosomal RNAs. RESULTS: We describe the graph theoretical identification and searching of highly hydrogen bonded base triples, where each base is involved in at least two hydrogen bonds with the other bases. Our approach correlates theoretically predicted base triples with literature-based compilations and other actual occurrences in crystal structures. The use of 'fuzzy' search tolerances has enabled us to discover a number of triple interaction types that have not been previously recorded nor predicted theoretically. CONCLUSIONS: Comparative analyses of different ribosomal RNA structures reveal several conserved base triple motifs in 50S rRNA structures, indicating an important role in structural stabilization and ultimately RNA function. Mohd Firdaus Raih, Anne-Marie Harrison, Peter Willett 0002, Peter J. Artymiuk |
BMC Bioinform. | 3 |
| 2010 | A Power-Enhanced Algorithm for Spatial Anomaly Detection in Binary Labelled Point Data Using the Spatial Scan Statistic
Simon Read, Peter A. Bath, Peter Willett 0002, Ravi Maheswaran |
KES (2) | 3 |
| 2005 | Research Paper: Use of Graph Theory to Identify Patterns of Deprivation and High Morbidity and Mortality in Public Health Data SetsabstractOBJECTIVE: An important part of public health is identifying patterns of poor health and deprivation. Specific patterns of poor health may be associated with features of the geographic environment where contamination or pollution may be occurring. For example, there may be clusters of poor health surrounding nuclear power stations, whereas major roads or rivers may be associated with areas of poor health alongside the feature in chains. Current methods are limited in their capacity to search for complex patterns in geographic data sets. The objective of this study was to determine whether graph theory could be used to identify patterns of geographic areas that have high levels of deprivation, morbidity, and mortality in a public health database. The geographic areas used in the study were enumeration districts (EDs), which are the lowest level of census geography in England and Wales, representing on average 200 households in the 1991 census. More specifically, the study aimed to identify chains of EDs with high deprivation, morbidity, and mortality that might be adjacent to specific types of geographic features, i.e., rivers or major roads. DESIGN: The maximum common subgraph (MCS) algorithm was used to search for seven query patterns of deprivation and poor health within the Trent region. Query pattern 1 represented a linear chain of five EDs and query patterns 2 to 7 represented the possible clusters of the five EDs. To identify chains of EDs with high deprivation, morbidity, and mortality, the results from the query patterns 2 to 7 were used to remove patterns (option 1) and EDs (option 2) from the results of query pattern 1. MEASUREMENTS: Data on the Townsend Material Deprivation Index, standardized long-term limiting illness and standardized all-cause mortality rates were used for the 10,665 EDs within the Trent region. RESULTS: The MCS algorithm retrieved a range of patterns and EDs from the database for the queries. Query pattern 1 identified 3,838 patterns containing a total of 195 EDs. When the patterns retrieved using query patterns 2 to 7 were removed from the 3,838 patterns using option 1, 1,704 patterns remained containing 161 EDs. When the EDs retrieved using query patterns 2 to 7 were removed from the 195 EDs identified by query pattern 1 using option 2, 12 EDs remained. The MCS algorithm was therefore able to reduce the numbers of patterns and EDs to allow manual examination for chains of EDs and for that which might be associated with them. CONCLUSION: The study demonstrates the potential of the MCS algorithm for searching for specific patterns of need. This method has potential for identifying such patterns in relation to local geographic features for public health. Peter A. Bath, Cheryl Craigs, Ravi Maheswaran, John W. Raymond, Peter Willett 0002 |
J. Am. Medical Informatics Assoc. | 5 |
| 2005 | Graph theoretic methods for the analysis of structural relationships in biological macromoleculesabstractAbstract Subgraph isomorphism and maximum common subgraph isomorphism algorithms from graph theory provide an effective and an efficient way of identifying structural relationships between biological macromolecules. They thus provide a natural complement to the pattern matching algorithms that are used in bioinformatics to identify sequence relationships. Examples are provided of the use of graph theory to analyze proteins for which three‐dimensional crystallographic or NMR structures are available, focusing on the use of the Bron‐Kerbosch clique detection algorithm to identify common folding motifs and of the Ullmann subgraph isomorphism algorithm to identify patterns of amino acid residues. Our methods are also applicable to other types of biological macromolecule, such as carbohydrate and nucleic acid structures. Peter J. Artymiuk, Ruth V. Spriggs, Peter Willett 0002 |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2004 | Chemoinformatics: an application domain for information retrieval techniquesabstractChemoinformatics is the generic name for the techniques used to represent, store and process information about the two-dimensional (2D) and three-dimensional (3D) structures of chemical molecules [1, 2]. Chemoinformatics has attracted much recent prominence as a result of developments in the methods that are used to synthesize new molecules and then to test them for biological activity. These developments have resulted in a massive increase in the amounts of structural and biological information that is available to support discovery programmes in the pharmaceutical and agrochemical industries.Chemoinformatics may appear to be far removed from information retrieval (IR), and there are indeed many significant differences, most notably in the use of graph representations to encode chemical molecules, rather than the strings that are used to encode text; however, there are also many similarities between the two fields, and this paper will exemplify some of these relationships. The most obvious area of similarity is in the principal types of database search that are carried out, with both application domains making extensive use of exact match, partial match and best match searching procedures: in the IR context these are known-item searching, Boolean searching and ranked-output searching; in the chemical context, these are structure searching, substructure searching and similarity searching. In IR, there is a natural distinction between an initial ranked-output search and one in which relevance feedback can be employed, where the keywords in the query statement are assigned weights based on their differential occurrences in known-relevant and known-nonrelevant documents. In the chemoinformatics technique called substructural analysis, substructural fragments are assigned weights based on their occurrence in molecules that do possess, and molecules that do not possess, some desired biological activity [3]. The analogy between relevance and biological activity has also resulted in the development of measures to quantify the effectiveness of chemical searching procedures that are based on the standard IR concepts of recall and precision [4].Analogies such as these have provided the basis for some of the chemoinformatics research carried out in Sheffield. The starting point was the recognition that techniques applicable to documents represented by keywords might also be applicable to molecules represented by substructural fragments. This led directly to the introduction of similarity searching, something that is now a standard tool in chemoinformatics software systems; in particular, its use for virtual screening, i.e., the ranking of a database in order of decreasing probability of activity so as to maximize the cost-effectiveness of biological testing [5]. Measures of inter-molecular structural similarity also lie at the heart of systems for clustering chemical databases: just as IR has the Cluster Hypothesis (similar documents tend to be relevant to the same requests) as a basis for document clustering, so the Similar Property Principle (similar molecules tend to have similar properties) has led to clustering becoming a well-established tool for the organization of large chemical databases [6]. More recently, we have applied another IR technique, the use of data fusion to combine different rankings of a database, to chemoinformatics and again found that it is equally applicable in this new domain [7].The many similarities between IR and chemoinformatics that have already been identified suggest that chemoinformatics is a domain of which IR researchers should be aware when considering the applicability of new techniques that they have developed. Peter Willett 0002 |
SIGIR | 1 |
| 2003 | Protein Structures and Information Extraction from Biological Texts: The PASTA SystemabstractMOTIVATION: The rapid increase in volume of protein structure literature means useful information may be hidden or lost in the published literature and the process of finding relevant material, sometimes the rate-determining factor in new research, may be arduous and slow. RESULTS: We describe the Protein Active Site Template Acquisition (PASTA) system, which addresses these problems by performing automatic extraction of information relating to the roles of specific amino acid residues in protein molecules from online scientific articles and abstracts. Both the terminology recognition and extraction capabilities of the system have been extensively evaluated against manually annotated data and the results compare favourably with state-of-the-art results obtained in less challenging domains. PASTA is the first information extraction (IE) system developed for the protein structure domain and one of the most thoroughly evaluated IE system operating on biological scientific text to date. AVAILABILITY: PASTA makes its extraction results available via a browser-based front end: http://www.dcs.shef.ac.uk/nlp/pasta/. The evaluation resources (manually annotated corpora) are also available through the website: http://www.dcs.shef.ac.uk/nlp/pasta/results.html. Robert J. Gaizauskas, George Demetriou, Peter J. Artymiuk, Peter Willett 0002 |
Bioinform. | 4 |
| 2002 | RASCAL: Calculation of Graph Similarity using Maximum Common Edge SubgraphsabstractA new graph similarity calculation procedure is introduced for comparing labeled graphs. Given a minimum similarity threshold, the procedure consists of an initial screening process to determine whether it is possible for the measure of similarity between the two graphs to exceed the minimum threshold, followed by a rigorous maximum common edge subgraph (MCES) detection algorithm to compute the exact degree and composition of similarity. The proposed MCES algorithm is based on a maximum clique formulation of the problem and is a significant improvement over other published algorithms. It presents new approaches to both lower and upper bounding as well as vertex selection. John W. Raymond, Eleanor J. Gardiner, Peter Willett 0002 |
Comput. J. | 3 |
| 1999 | Computational analysis of molecular diversity for drug discoveryabstractThis paper provides a brief introduction to modem techniques for the computational analysis of molecular diversity and then discusses current work on the selection of structurally diverse sets of compounds using dissimilarity-based and partition-based approaches to selection.The first study involves a genetic algorithm for designing combinatorial libraries; the second study involves four schemes for defining the sizes of the bins in a partition, and discusses ways in which the effectiveness of such schemes can be measured. Martin J. Bayley, Valerie J. Gillet, Peter Willett 0002, John Bradshaw, Darren V. S. Green |
RECOMB | 3 |
| 1998 | Similarity and Dissimilarity Methods for Processing Chemical Structure DatabasesabstractThis paper reviews measures of similarity and dissimilarity between pairs of chemical molecules and the use of such measures for processing chemical databases. The applications discussed include similarity searching, database clustering and diversity analysis, focusing upon measures that are based on fragment bit-string occurrence data. The paper then discusses recent work on the calculation of similarity by aligning molecular fields and on the selection of structurally diverse subsets of chemical databases. Valerie J. Gillet, David J. Wild 0002, Peter Willett 0002, John Bradshaw |
Comput. J. | 3 |
| 1996 | On the Creation of Hypertext Links in Full-Text Documents: Measurement of Retrieval EffectivenessabstractAn important stage in the process of retrieval of objects from a hypertext database is the creation of a set of internodal links that are intended to represent the relationships existing between objects; this operation is often undertaken manually, just as index terms are often manually assigned to documents in a conventional retrieval system. In an earlier article (Ellis, D., Furner-Hines, J., & Willett, P., 1994b), the results were published of a study in which several different sets of links were inserted, each by a different person, between the paragraphs of each of a number of full-text documents. These results showed little similarity between the link-sets, a finding that was comparable with those of studies of inter-indexer consistency, which suggest that there is generally only a low level of agreement between the sets of index terms assigned to a document by different indexers. In this article, a description is provided of an investigation into the nature of the relationship existing between (i) the levels of inter-linker consistency obtaining among the group of hypertext databases used in our earlier experiments, and (ii) the levels of effectiveness of a number of searches carried out in those databases. An account is given of the implementation of the searches and of the methods used in the calculation of numerical values expressing their effectiveness. Analysis of the results of a comparison between recorded levels of consistency and those of effectiveness does not allow us to draw conclusions about the consistency-effectiveness relationship that are equivalent to those drawn in comparable studies of inter-indexer consistency. © 1996 John Wiley & Sons, Inc. David Ellis, Jonathan Furner, Peter Willett 0002 |
J. Am. Soc. Inf. Sci. | 3 |
| 1994 | On the Measurement of Inter-Linker Consistency and Retrieval Effectiveness in Hypertext Databases
David Ellis, Jonathan Furner, Peter Willett 0002 |
SIGIR | 3 |
| 1993 | On the Non-Random Nature of the Nearest-Neighbour Document Clusters
Rachel J. Shaw, Peter Willett 0002 |
Inf. Process. Manag. | 2 |
| 1992 | Searching for Historical Word-Forms in a Database of 17th-Century English Text Using Spelling-Correction MethodsabstractThis paper discusses the application of algorithmic spelling-correction techniques to the identification of those words in a database of 17th century English text that are most similar to a query word in modern English. The experiments have used n-gram matching, non-phonetic coding and dynamic programming methods for spelling correction, and have demonstrated that high-recall searches can be carried out, although some of the searches are very demanding of computational resources. The methods are, in principle, applicable to historical texts in many languages and from many diffeent periods. Alexander M. Robertson, Peter Willett 0002 |
SIGIR | 2 |
| 1992 | The Effectiveness of Stemming for Natural-Language Access to Slovene Textual DataabstractThere have been several studies of the use of stemming algorithms for conflating morphological variants in free-text retrieval systems. Comparison of stemmed and nonconflated searches suggests that there are no significant increases in the effectiveness of retrieval when stemming is applied to English-language documents and queries. This article reports the use of stemming on Slovene-language documents and queries, and demonstrates that the use of an appropriate stemming algorithm results in a large, and statistically significant, increase in retrieval effectiveness when compared with nonconflated processing; similar comments apply to the use of manual, right-hand truncation. A comparison is made with stemming of English versions of the same documents and queries and it is concluded that the effectiveness of a stemming algorithm is determined by the morphological complexity of the language that it is designed to process. © 1992 John Wiley & Sons, Inc. Mirko Popovic, Peter Willett 0002 |
J. Am. Soc. Inf. Sci. | 2 |
| 1991 | Network design for the implementation of text searching using a multicomputer
Janey K. Cringean, Roger England, Gordon A. Manson, Peter Willett 0002 |
Inf. Process. Manag. | 4 |
| 1991 | The limitations of term co-occurrence data for query expansion in document retrieval systemsabstractTerm cooccurrence data has been extensively used in document retrieval systems for the identification of indexing terms that are similar to those that have been specified in a user query: these similar terms can then be used to augment the original query statement. Despite the plausibility of this approach to query expansion, the retrieval effectiveness of the expanded queries is often no greater than, or even less than, the effectiveness of the unexpanded queries. This article demonstrates that the similar terms identified by cooccurrence data in a query expansion system tend to occur very frequently in the database that is being searched. Unfortunately, frequent terms tend to discriminate poorly between relevant and nonrelevant documents, and the general effect of query expansion is thus to add terms that do little or nothing to improve the discriminatory power of the original query. © 1991 John Wiley & Sons, Inc. Helen J. Peat, Peter Willett 0002 |
J. Am. Soc. Inf. Sci. | 2 |
| 1990 | Parallel Text Searching in Serial Files Using a Processor FarmabstractThis paper discusses the implementation of a parallel text retrieval system using a microprocessor network. The system is designed to allow fast searching in document databases organised using the serial file structure, with a very rapid initial text signature search being followed by a more detailed, but more time-consuming, pattern matching search. The network is built from transputers, high performance microprocessors developed specifically for the construction of highly parallel computing systems, which are linked together in a processor farm. The paper discusses the design and implementation of processor farms, and then reports our initial studies of the efficiency of searching that can be achieved using this approach to text retrieval from serial files. Janey K. Cringean, Roger England, Gordon A. Manson, Peter Willett 0002 |
SIGIR | 4 |
| 1989 | Comparison of Hierarchie Agglomerative Clustering Methods for Document RetrievalabstractThis paper considers the use of the single linkage, complete linkage, group average and Ward hierarchic agglomerative clustering methods for document retrieval. The methods are used to cluster seven document test collections for which queries and relevance judgements are available. Several retrieval strategies are described which allow searches to be carried out of the clustered document files resulting from the use of the four methods. These searches suggest that the group average method is the most suitable for document clustering purposes; however, searches of the unclustered document collections and of a simpler type of clustered file (based on pairs of nearest neighbours) usually result in better levels of retrieval effectiveness than searches of the clustered collections. Abdelmoula El-Hamdouchi, Peter Willett 0002 |
Comput. J. | 2 |
| 1989 | Retrieving documents by plausible inference: An experimental study
W. Bruce Croft, T. J. Lucia, Janey K. Cringean, Peter Willett 0002 |
Inf. Process. Manag. | 4 |
| 1988 | An improved algorithm for the calculation of exact term discrimination values
Abdelmoula El-Hamdouchi, Peter Willett 0002 |
Inf. Process. Manag. | 2 |
| 1988 | Recent trends in hierarchic document clustering: A critical review
Peter Willett 0002 |
Inf. Process. Manag. | 1 |
| 1988 | Bibliographic pattern matching using the ICL Distributed Array ProcessorabstractThis article discusses the use of the ICL Distributed Array Processor (DAP), a highly parallel array processor containing 4096 processing elements, for pattern-matching operations in a bibliographic retrieval system. The hardware and software features of the DAP are described, together with a pattern-matching algorithm that makes full use of the DAP's parallelism. This algorithm can be used to search for exact patterns, right-hand or left-hand truncated patterns, embedded patterns, patterns containing fixed-length (FLDC) or variable-length (VLDC) don't care characters, and patterns that specify adjacent words. Experiments with a set of 898 query patterns and 2472 titles and abstracts from the INSPEC database show that right-hand truncation searches can be carried out on the DAP about ten times faster than searches using the Aho-Corasick pattern-matching algorithm on a Prime 550 minicomputer, the relative advantage of the DAP increasing linearly with a decrease in the number of query patterns. Simulated FLDC and VLDC searches on the DAP take about two and three times as long as right-hand truncation searches, respectively. © 1988 John Wiley & Sons, Inc. David M. Carroll, Christine A. Pogue, Peter Willett 0002 |
J. Am. Soc. Inf. Sci. | 3 |
| 1988 | Chemical graph matching using transputer networks
Andrew T. Brint, Valerie J. Gillet, Michael F. Lynch, Peter Willett 0002, Gordon A. Manson, George A. Wilson |
Parallel Comput. | 4 |
| 1988 | Searching and clustering of databases using the ICL distributed array processor
Christine A. Pogue, Edie M. Rasmussen, Peter Willett 0002 |
Parallel Comput. | 3 |
| 1987 | Non-Hierarchic Document Clustering Using the ICL Distributed Array ProcessorabstractThis paper considers the suitability and efficiency of a highly parallel computer, the ICL Distributed Array Processor (DAP), for document clustering. Algorithms are described for the implementation of the single-pass and reallocation clustering methods on the DAP and on a conventional mainframe computer. These methods are used to classify the Cranfield, Vaswani and UKCIS document test collections. The results suggest that the parallel architecture of the DAP is not well suited to the variable-length records which characterise bibliographic data. Edie M. Rasmussen, Peter Willett 0002 |
SIGIR | 2 |
| 1987 | Current research into chemical and textual information retrieval at the department of information studies, University of Sheffield
Michael F. Lynch, Peter Willett 0002 |
Inf. Process. Manag. | 2 |
| 1987 | Use of text signatures for document retrieval in a highly parallel environment
Christine A. Pogue, Peter Willett 0002 |
Parallel Comput. | 2 |
| 1986 | Hierarchic Document Clustering Using Ward's MethodabstractIn this paper, we discuss the application of a recent hierarchic clustering algorithm to the automatic classification of files of documents. Whereas most hierarchic clustering algorithms involve the generation and updating of an inter-object dissimilarity matrix, this new algorithm is based upon a series of nearest neighbor searches. Such an approach is appropriate to several clustering methods, including Ward's method which has been shown to perform well in experimental studies of hierarchic document clustering. A description is given of heuristics which can increase the efficiency of the new algorithm when it is used to cluster three document collections by Ward's method. Abdelmoula El-Hamdouchi, Peter Willett 0002 |
SIGIR | 2 |
| 1986 | Compression of nucleic acid and protein sequence dataabstractThis paper describes the application of text compression methods to machine-readable files of nucleic acid and protein sequence data. Two main methods are used to reduce the storage requirements of such files, these being n-gram coding and run-length coding. A Pascal program combining both of these techniques resulted in a compression figure of 74.6% for the GenBank data-base and a program that used only n-gram coding gave a compression figure of 42.8% for the Protein Identification Resource database. J. Richard Walker, Peter Willett 0002 |
Comput. Appl. Biosci. | 2 |
| 1986 | Criteria for the Selection of Search Strategies in Best-Match Document-Retrieval Systems
Fiona M. McCall, Peter Willett 0002 |
Int. J. Man Mach. Stud. | 2 |
| 1986 | Using interdocument similarity information in document retrieval systemsabstractThe first part of this paper reports a comparative study of the document classifications produced by the use of the single linkage, complete linkage, group average, and Ward clustering methods. Studies of cluster member-ship and of the effectiveness of cluster searches support previous findings that suggest that the single linkage classifications are rather different from those produced by the other three methods. These latter methods all pro-duce large numbers of small clusters containing just pairs of documents. This finding motivates the work reported in the second part of the paper, which considers the use of clusters consisting of a document together with that document with which it is most similar. A com-parison of the use of such clusters with conventional best match searches using seven document test collec-tions suggests that the two types of search are of com-parable effectiveness, but they retrieve noticeably differ-ent sets of relevant documents. Alan Griffiths, H. Claire Luckhurst, Peter Willett 0002 |
J. Am. Soc. Inf. Sci. | 3 |
| 1985 | An algorithm for the calculation of exact term discrimination values
Peter Willett 0002 |
Inf. Process. Manag. | 1 |
| 1984 | A note on the use of nearest neighbors for implementing single linkage document classificationsabstractAbstract Best match search algorithms provide an efficient means of identifying the sets of nearest neighbors for each of the documents in a collection. These sets contain much of the important similarity data contained in a full interdocument similarity matrix and may be used for the generation of hierarchic document classifications, such as those arising from the use of the single linkage clustering method. Cluster based retrieval experiments based upon such classifications are shown to give results that are comparable in effectiveness with those obtained using the full similarity matrix. Peter Willett 0002 |
J. Am. Soc. Inf. Sci. | 1 |
| 1983 | Automatic Spelling Correction Using a Trigram Similarity Measure
Richard C. Angell, George E. Freund, Peter Willett 0002 |
Inf. Process. Manag. | 3 |
| 1983 | Document Retrieval Using a Serial Bit String Search
Alan F. Harding, Michael F. Lynch, Peter Willett 0002 |
Inf. Process. Manag. | 3 |
| 1982 | The effect of subject matter on the automatic indexing of full textabstractAbstract A recently suggested method for the automatic indexing of full text is applied to extracts from the Brown Corpus. Scientific and technological extracts are found to give rise to a much larger number of index terms than humanities and social science extracts. These results would appear to arise from differences in the word frequency distributions for each type of subject. Mary E. Rowbottom, Peter Willett 0002 |
J. Am. Soc. Inf. Sci. | 2 |
| 1982 | Some Current Information Retrieval Research in the United Kingdom
Peter Willett 0002 |
J. Am. Soc. Inf. Sci. | 1 |
| 1981 | A fast procedure for the calculation of similarity coefficients in automatic classification
Peter Willett 0002 |
Inf. Process. Manag. | 1 |
| 1980 | Indexing exhaustivity and the computation of similarity matricesabstractAbstract Some of the automatic classification procedures used in information retrieval derive clusters of documents from an intermediate similarity matrix, the computation of which involves comparing each of the documents in the collection with all of the others. It has recently been suggested that many of these comparisons, specifically those between documents having no terms in common, may be avoided by means of the use of an inverted file to the document collection. This communication shows that the approach will effect reductions in the number of interdocument comparisons only if the documents are each indexed by a limited number of indexing terms; if exhaustive indexing is used, many document pairs will be compared several times over and the computation will be greater than when conventional approaches are used to generate the similarity matrix. Alan F. Harding, Peter Willett 0002 |
J. Am. Soc. Inf. Sci. | 2 |