VLDB 2026 Research / reviewers in the wild / expert
Wellington Santos Martins
dblp:120/9618 · also Wellington Martins, Wellington S. Martins
· DBLP profile ↗
16ranked-venue papers
1as first author
2since 2021 · last 2021
0000-0002-9641-2565ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 1 since 2021Artificial intelligence and machine learning · 6Systems, architecture and hardware · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 54% Data mining · 46% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 50% Parallel and multicore computing · 50% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › ranking › learning to rank
feature selection for ranking |
0.4 | 1 | 2019 | Risk-Sensitive Learning to Rank with Evolutionary Multi-Objective Feature Selection · ACM Trans. Inf. Syst. 2019 |
Information retrieval › ranking
learning to rank |
0.4 | 1 | 2019 | Risk-Sensitive Learning to Rank with Evolutionary Multi-Objective Feature Selection · ACM Trans. Inf. Syst. 2019 |
Data mining › predictive modeling
classification |
0.2 | 1 | 2015 | An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel Computing · SIGIR 2015 |
Data mining › predictive modeling › classification › nearest neighbor classification
k-nearest neighbor classification |
0.2 | 1 | 2015 | An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel Computing · SIGIR 2015 |
Data mining › text mining
text classification |
0.2 | 1 | 2015 | An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel Computing · SIGIR 2015 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2015 | An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel Computing · SIGIR 2015 |
Parallel and multicore computing › parallel algorithms
massively parallel algorithms |
0.1 | 1 | 2015 | An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel Computing · SIGIR 2015 |
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis |
0.0 | 1 | 2002 | TROLL-Tandem Repeat Occurrence Locator · Bioinform. 2002 |
Bioinformatics and computational biology › sequence analysis › repeat detection
tandem repeat detection |
0.0 | 1 | 2002 | TROLL-Tandem Repeat Occurrence Locator · Bioinform. 2002 |
Algorithms and data structures › sequence algorithms › string algorithms › string matching
aho-corasick algorithm |
0.0 | 1 | 2002 | TROLL-Tandem Repeat Occurrence Locator · Bioinform. 2002 |
Algorithms and data structures › sequence algorithms › string algorithms
string matching |
0.0 | 1 | 2002 | TROLL-Tandem Repeat Occurrence Locator · Bioinform. 2002 |
Methods — techniques the papers use, named apart from their topics
parallel kNN · 0.4GPU acceleration · 0.4multi-objective optimization · 0.4evolutionary algorithm · 0.4SPEA2 · 0.4aho-corasick algorithm · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Computer architecture and high performance computingabstractIn this special issue of Concurrency and Computation Practice and Experience, we are pleased to present eight selected papers that were previously presented at the Brazilian "XX Simpósio em Sistemas Computacionais de Alto Desempenho," WSCAD 2019. The event was held in conjunction with the 31st International Symposium on Computer Architecture and High Performance Computing, SBAC-PAD 2019, in Campo Grande, MS, Brazil, from October 15 to 18, 2019. The WSCAD workshop has been presenting important research in the fields of computer architectures, high performance computing, and distributed systems, since the beginning of the 2000s. The scope of the current special issue is broad and representative, with different forms of contributions to our discipline: methodological papers, technology papers, application papers, and system papers. The topics covered in the papers include architecture issues, compiler optimization, performance evaluation, parallel algorithms, energy efficiency, and applications. The title of the first paper is "Structural testing for communication events into loops of message-passing parallel programs," by Diaz et al.1 In this paper, the authors propose new structural testing criteria for message-passing parallel programs, focusing on defects from communication primitives into loops. A new test model is presented to support their criteria for structural testing of MPI-applications. The testing criteria are validated through experimental studies using a tool called ValiMPI. The results show that unknown defects from communication and synchronization events can be revealed in different loop iterations, increasing the quality of message-passing parallel programs. In the second contribution, entitled "Smart selection of optimizations in dynamic compilers," Rosario et al.2 present an approach that uses machine learning to select sequences of optimization for dynamic compilation that considers both code quality and compilation overhead. Their approach starts by training a model, offline, with a knowledge bank of those sequences with low overhead and high-quality code generation capability using a genetic heuristic. Then, this bank is used to guide the smart selection of optimizations sequences for the compilation of code fragments during the emulation of an application. The proposed strategy is evaluated in two LLVM-based dynamic binary translators, namely, OI-DBT and HQEMU, showing that these two translators can achieve average speedups of 1.26× and 1.15× in MiBench and Spec Cpu benchmarks, respectively. In the third contribution, entitled "Memory allocation anomalies in high-performance computing applications: A study with numerical simulations," Gomes et al.3 propose a method for identifying, locating, characterizing, and fixing allocation anomalies, and a tool for developers to apply the method. A numerical simulator that approximates the solutions to partial differential equations using a finite element method is used in the experiments. It is shown that taming allocation anomalies in the simulator reduces both its execution time and the memory footprint of its processes, irrespective of the specific heap allocator being employed with it. They conclude that the developer of HPC applications can benefit from the method and tool during the software development cycle. The fourth contribution, entitled "Investigating memory prefetcher performance over parallel applications: From real to simulated," by Girelli et al.,4 contributes to shed light on the memory prefetcher's role in the performance of parallel high-performance computing applications, considering the prefetcher algorithms offered by both the real hardware and the simulators. The authors performed a careful experimental investigation, executing the NAS parallel benchmark (NPB) on a real Skylake machine, and as well in a simulated environment with the ZSim and Sniper simulators, taking into account the prefetcher algorithms offered by both Skylake and the simulators. The experimental results show that: (i) prefetching from the L3 to L2 cache presents better performance gains, (ii) the memory contention in the parallel execution constrains the prefetcher's effect, (iii) Skylake's parallel memory contention is poorly simulated by ZSim and Sniper, and (iv) Skylake's noninclusive L3 cache hinders the accurate simulation of NPB with the Sniper's prefetchers. In the fifth contribution, entitled "Energy efficiency and portability of oil and gas simulations on multicore and graphics processing unit architectures," Serpa et al.5 propose three optimizations for an oil and gas application, reverse time migration (RTM), which reduce the floating-point operations by changing the equation derivatives. They evaluate these optimizations in different multicore and GPU architectures, investigating the impact of different APIs on the performance, energy efficiency, and portability of the code. The experimental results show that the dedicated CUDA implementation running on the NVIDIA Volta architecture has the best performance and energy efficiency for RTM on GPUs, while the OpenMP version is the best for Intel Broadwell in the multicore. Also, the OpenACC version, which has a lower programming effort and executes on both architectures, has up to 20% better performance and energy efficiency than the nonportable ones. In the sixth paper, entitled "An open computing language-based parallel Brute Force algorithm for formal concept analysis on heterogeneous architectures," Novais et al.6 propose and evaluate an Open Computing Language (OpenCL)-based Brute Force algorithm for formal concept extraction on heterogeneous architectures (CPU + GPU and CPU + FPGA). The CPU + GPU architecture presents higher performance and scalability than other architectures when the Brute Force algorithm processes high dimensional contexts with many objects and attributes. Their parallel approach shows performance results up to 18× better than a smarter sequential algorithm called Data-Peeler. Moreover, the Brute Force algorithm running on CPU + GPU architecture has greater energy efficiency, reaching at least 1.79× more operations per energy consumption than other algorithms on different architectures explored in the work. In the seventh paper, entitled "Contextual contracts for component-oriented resource abstraction in a cloud of high performance computing services," Junior et al.7 present HPC Shelf, a cloud computing services platform to build and deploy large-scale parallel computing systems. They introduce Alite, the contextual contract system of HPC Shelf, to select component implementations according to requirements of the host application, target parallel computing platform characteristics (e.g., clusters and MPPs), quality of service (QoS) properties, and cost restrictions. It is evaluated through a small-scale case study employing two complementary component-based frameworks. The first one aims to represent components that implement linear algebra computations based on the BLAS interface. In turn, the second one aims to represent parallel computing platforms on the IaaS cloud offered by Amazon EC2 Service. The last paper in this special issue, "High-performance IO for seismic processing on the cloud" authored by Guimarães et al.,8 analyzes the main file structures currently used to store seismic data and propose a new intermediate data structure to improve IO performance while still complying with established standards. They show that, throughout a common workflow in seismic data analysis, the IO performance gain greatly surpasses the overhead of translating data to the intermediate structure. The approach enables a speedup of up to 208 times in reading time when using classical standards (e.g., SEG-Y) and the intermediate structure is up to 1.8 times more efficient than modern formats (e.g., ASDF). Considering cache-friendly applications, the speedups over the direct use of SEG-Y reach 8000 times. They also performed a cost analysis on the AWS cloud showing that HDDs can be 1.25 times more cost-effective than SSDs. The research papers presented in this special issue provide insights in fields related to high performance computing, including performance evaluation, parallel algorithms, and applications in science and engineering. We believe that the main contributions presented in the research papers are timely and important, and hope that readers can benefit from the papers and contribute to these rapidly growing areas. Many individuals contributed a great deal of time and energy toward the success of this special issue. We would like to thank all the authors who provided valuable contributions to this special issue. We are also grateful to the reviewers for their many hours of dedicated efforts, with valuable feedback to the authors. Finally, we would also like to express our gratitude to the Editor-in-Chief of CCPE, for his advice, vision, and support, making this special issue possible. Raphael Y. de Camargo, Fabrizio Marozzo, Wellington Santos Martins |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | On the cost-effectiveness of neural and non-neural approaches and representations for text classification: A comprehensive comparative study
Washington Cunha, Vítor Mangaravite, Christian Gomes, Sérgio D. Canuto, Elaine Resende, Cecilia Nascimento, Felipe Viegas, Celso França, Wellington Santos Martins, Jussara M. Almeida, Thierson Couto, Leonardo Rocha 0001, Marcos André Gonçalves |
Inf. Process. Manag. | 9 |
| 2020 | "Keep it Simple, Lazy" - MetaLazy: A New MetaStrategy for Lazy Text ClassificationabstractRecent advances in text-related tasks on the Web, such as text (topic) classification and sentiment analysis, have been made possible by exploiting mostly the "rule of more": more data (massive amounts) more computing power, more complex solutions. We propose a shift in the paradigm to do "more with less" by focusing, at maximum extent, just on the task at hand (e.g., classify a single test instance). Accordingly, we propose MetaLazy, a new supervised lazy text classification meta-strategy that greatly extends the scope of lazy solutions. Lazy classifiers postpone the creation of a classification model until a given test instance for decision making is given. MetaLazy exploits new ideas and solutions, which have in common their lazy nature, producing altogether a solution for text classification, which is simpler, more efficient, and less data demanding than new alternatives. It extends and evolves the lazy creation of the model for the test instance by allowing: (i) to dynamically choose the best classifier for the task; (ii) the exploration of distances in the neighborhood of the test document when learning a classification model, thus diminishing the importance of irrelevant training instances; and (iii) a better representational space for training and test documents by augmenting them, in a lazy fashion, with new co-occurrence based features considering just those observed in the specific test instance. In a sizeable experimental evaluation, considering topics and sentiment analysis datasets and nine baselines, we show that our MetaLazy instantiations are among the top performers in most situations, even when compared to state-of-the-art deep learning classifiers such as Deep Network Transformer Architectures. Luiz Felipe Mendes, Marcos André Gonçalves, Washington Cunha, Leonardo Rocha 0001, Thierson Couto, Wellington Santos Martins |
CIKM | 6 |
| 2019 | Parallel rule-based selective sampling and on-demand learning to rankabstractSummary Learning to rank (L2R) works by constructing a ranking model from training data so that, given a new query, the model is able to generate an effective rank of the objects for the query. Almost all work in L2R focus on ranking accuracy leaving performance and scalability overlooked. However, performance is a critical factor, especially when dealing with on‐demand queries. In this scenario, Learning to Rank using association rules has been shown to be extremely effective but only at a high computational cost. In this work, we show how to exploit parallelism on rule‐based systems to: i) drastically reduce L2R training datasets using selective sampling and ii) to generate query customized ranking models on the fly. We present parallel algorithms and GPU implementations for these two tasks showing that dataset reduction takes only a few seconds with speedups up to 148x over a serial baseline, and that queries can be processed in only a few milliseconds with speedups of 1000x over a serial baseline and 29x over a parallel baseline for the best case. We also extend the implementations to work with multiple GPUs, further increasing the speedup over the baselines and showing the scalability of our proposed algorithms. Mateus Ferreira e Freitas, Daniel Xavier de Sousa, Wellington Santos Martins, Thierson Couto, Rodrigo M. Silva, Marcos André Gonçalves |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Risk-Sensitive Learning to Rank with Evolutionary Multi-Objective Feature SelectionabstractLearning to Rank (L2R) is one of the main research lines in Information Retrieval. Risk-sensitive L2R is a sub-area of L2R that tries to learn models that are good on average while at the same time reducing the risk of performing poorly in a few but important queries (e.g., medical or legal queries). One way of reducing risk in learned models is by selecting and removing noisy, redundant features, or features that promote some queries to the detriment of others. This is exacerbated by learning methods that usually maximize an average metric (e.g., mean average precision (MAP) or Normalized Discounted Cumulative Gain (NDCG)). However, historically, feature selection (FS) methods have focused only on effectiveness and feature reduction as the main objectives. Accordingly, in this work, we propose to evaluate FS for L2R with an additional objective in mind, namely risk-sensitiveness . We present novel single and multi-objective criteria to optimize feature reduction, effectiveness, and risk-sensitiveness, all at the same time. We also introduce a new methodology to explore the search space, suggesting effective and efficient extensions of a well-known Evolutionary Algorithm (SPEA2) for FS applied to L2R. Our experiments show that explicitly including risk as an objective criterion is crucial to achieving a more effective and risk-sensitive performance. We also provide a thorough analysis of our methodology and experimental results. Daniel Xavier de Sousa, Sérgio D. Canuto, Marcos André Gonçalves, Thierson Couto, Wellington Santos Martins |
ACM Trans. Inf. Syst. | 5 |
| 2018 | A New Word Embedding Approach to Evaluate Potential Fixes for Automated Program RepairabstractDebugging is frequently a manual and costly task. Some work recently presented automated program repair methods aiming to reduce debugging time. Despite different approaches, the process of evaluating source code patches (potential fixes) is crucial for most of them, e.g., generate-and-validate systems. The evaluation is a complex task given that patches with different syntaxes might share the same semantics, behaving equally for typically limited specifications. Hence, many approaches fail to better explore the search space of patches, leading them to not reach the patch that fixes the bug. Some research points that a buggy code is more entropic, i.e., less natural than its fixed version. So, this work proposes applying Word2vec, a word embedding model, to improve the repair evaluation process based on the naturalness obtained from a corpus of known fixes. Word2vec captures co-occurrence relationships between words in a given context and then predicts the contextual words of a given word. This technique has been applied to deal with richer semantic relationships in a text. Word2vec evaluates patches according to distances of document vectors and a softmax output layer. We analyze the performance of our proposal with mutated patches created from correct source codes and we simulate potential fixes generated by automated program repair approaches. Thus, the main contribution of this paper is a new method to evaluate patches used in automated program repair methods. The results show that Word2vec-based metrics are capable of analyzing source code naturalness and be used to evaluate source code patches. Leonardo Afonso Amorim, Mateus F. Freitas, Altino Dantas, Eduardo Faria de Souza, Celso G. Camilo-Junior, Wellington Santos Martins |
IJCNN | 6 |
| 2018 | Exploiting efficient and effective lazy Semi-Bayesian strategies for text classification
Felipe Viegas, Leonardo Rocha 0001, Elaine Resende, Thiago Salles, Wellington Santos Martins, Mateus Ferreira e Freitas, Marcos André Gonçalves |
Neurocomputing | 5 |
| 2016 | Incorporating Risk-Sensitiveness into Feature Selection for Learning to RankabstractLearning to Rank (L2R) is currently an essential task in basically all types of information systems given the huge and ever increasing amount of data made available. While many solutions have been proposed to improve L2R functions, relatively little attention has been paid to the task of improving the quality of the feature space. L2R strategies usually rely on dense feature representations, which contain noisy or redundant features, increasing the cost of the learning process, without any benefits. Although feature selection (FS) strategies can be applied to reduce dimensionality and noise, side effects of such procedures have been neglected, such as the risk of getting very poor predictions in a few (but important) queries. In this paper we propose multi-objective FS strategies that optimize both aspects at the same time: ranking performance and risk-sensitive evaluation. For this, we approximate the Pareto-optimal set for multi-objective optimization in a new and original application to L2R. Our contributions include novel FS methods for L2R which optimize multiple, potentially conflicting, criteria. In particular, one of the objectives (risk-sensitive evaluation) has never been optimized in the context of FS for L2R before. Our experimental evaluation shows that our proposed methods select features that are more effective (ranking performance) and low-risk than those selected by other state-of-the-art FS methods. Daniel Xavier de Sousa, Sérgio D. Canuto, Thierson Couto, Wellington Santos Martins, Marcos André Gonçalves |
CIKM | 4 |
| 2015 | Parallel Lazy Semi-Naive Bayes Strategies for Effective and Efficient Document ClassificationabstractAutomatic Document Classification (ADC) is the basis of many important applications such as spam filtering and content organization. Naive Bayes (NB) approaches are a widely used classification paradigm, due to their simplicity, efficiency, absence of parameters and effectiveness. However, they do not present competitive effectiveness when compared to other modern statistical learning methods, such as SVMs. This is related to some characteristics of real document collections, such as class imbalance, feature sparseness and strong relationships among attributes. In this paper, we investigate whether the relaxation of the NB feature independence assumption (aka, Semi-NB approaches) can improve its effectiveness in large text collections. We propose four new Lazy Semi-NB strategies that exploit different ideas for alleviating the NB independence assumption. By being lazy, our solutions focus only on the most important features to classify a given test document, overcoming some Semi-NB issues when applied to ADC such as bias towards larger classes and overfitting and/or lack of generalization of the models. We demonstrate that our Lazy Semi-NB proposals can produce superior effectiveness when compared to state-of-the-art ADC classifiers such as SVM and KNN. Moreover, to overcome some efficiency issues of combining Semi-NB and lazy strategies, we take advantage of current manycore GPU architectures and present a massively parallelized version of the Semi-NB approaches. Our experimental results show that speedups of up to 63.36 times can be obtained when compared to serial solutions, making our proposals very practical in real-situations. Felipe Viegas, Marcos André Gonçalves, Wellington Santos Martins, Leonardo Rocha 0001 |
CIKM | 3 |
| 2015 | An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel ComputingabstractThe unprecedented growth of available data nowadays has stimulated the development of new methods for organizing and extracting useful knowledge from this immense amount of data. Automatic Document Classification (ADC) is one of such methods, that uses machine learning techniques to build models capable of automatically associating documents to well-defined semantic classes. ADC is the basis of many important applications such as language identification, sentiment analysis, recommender systems, spam filtering, among others. Recently, the use of meta-features has been shown to substantially improve the effectiveness of ADC algorithms. In particular, the use of meta-features that make a combined use of local information (through kNN-based features) and global information (through category centroids) has produced promising results. However, the generation of these meta-features is very costly in terms of both, memory consumption and runtime since there is the need to constantly call the kNN algorithm. We take advantage of the current manycore GPU architecture and present a massively parallel version of the kNN algorithm for highly dimensional and sparse datasets (which is the case for ADC). Our experimental results show that we can obtain speedup gains of up to 15x while reducing memory consumption in more than 5000x when compared to a state-of-the-art parallel baseline. This opens up the possibility of applying meta-features based classification in large collections of documents, that would otherwise take too much time or require the use of an expensive computational platform. Sérgio D. Canuto, Marcos André Gonçalves, Wisllay M. V. dos Santos, Thierson Couto, Wellington Santos Martins |
SIGIR | 5 |
| 2015 | A variant of k-nearest neighbors search with cyclically permuted query points for rotation-invariant image processing
Les R. Foulds, Jorge P. de Morais Neto, Humberto J. Longo, Hugo A. D. do Nascimento, Wellington Santos Martins |
Discret. Appl. Math. | 5 |
| 2014 | On Efficient Meta-Level Features for Effective Text ClassificationabstractThis paper addresses the problem of automatically learning to classify texts by exploiting information derived from meta-level features (i.e., features derived from the original bag-of-words representation). We propose new meta-level features derived from the class distribution, the entropy and the within-class cohesion observed in the k nearest neighbors of a given test document x, as well as from the distribution of distances of x to these neighbors. The set of proposed features is capable of transforming the original feature space into a new one, potentially smaller and more informed. Experiments performed with several standard datasets demonstrate that the effectiveness of the proposed meta-level features is not only much superior than the traditional bag-of-word representation but also superior to other state-of-art meta-level features previously proposed in the literature. Moreover, the proposed meta-features can be computed about three times faster than the existing meta-level ones, making our proposal much more scalable. We also demonstrate that the combination of our meta features and the original set of features produce significant improvements when compared to each feature set used in isolation. Sérgio D. Canuto, Thiago Salles, Marcos André Gonçalves, Leonardo Rocha 0001, Gabriel Spada Ramos, Luiz Gonçalves 0001, Thierson Couto, Wellington Santos Martins |
CIKM | 8 |
| 2013 | SUNPLIN: Simulation with Uncertainty for Phylogenetic InvestigationsabstractBACKGROUND: Phylogenetic comparative analyses usually rely on a single consensus phylogenetic tree in order to study evolutionary processes. However, most phylogenetic trees are incomplete with regard to species sampling, which may critically compromise analyses. Some approaches have been proposed to integrate non-molecular phylogenetic information into incomplete molecular phylogenies. An expanded tree approach consists of adding missing species to random locations within their clade. The information contained in the topology of the resulting expanded trees can be captured by the pairwise phylogenetic distance between species and stored in a matrix for further statistical analysis. Thus, the random expansion and processing of multiple phylogenetic trees can be used to estimate the phylogenetic uncertainty through a simulation procedure. Because of the computational burden required, unless this procedure is efficiently implemented, the analyses are of limited applicability. RESULTS: In this paper, we present efficient algorithms and implementations for randomly expanding and processing phylogenetic trees so that simulations involved in comparative phylogenetic analysis with uncertainty can be conducted in a reasonable time. We propose algorithms for both randomly expanding trees and calculating distance matrices. We made available the source code, which was written in the C++ language. The code may be used as a standalone program or as a shared object in the R system. The software can also be used as a web service through the link: http://purl.oclc.org/NET/sunplin/. CONCLUSION: We compare our implementations to similar solutions and show that significant performance gains can be obtained. Our results open up the possibility of accounting for phylogenetic uncertainty in evolutionary and ecological analyses of large datasets. Wellington Santos Martins, Welton Couto Carmo, Humberto J. Longo, Thierson Couto, Thiago Fernando Rangel |
BMC Bioinform. | 1 |
| 2012 | Improving On-Demand Learning to Rank through Parallelism
Daniel Xavier de Sousa, Thierson Couto, Wellington Santos Martins, Rodrigo M. Silva, Marcos André Gonçalves |
WISE | 3 |
| 2002 | TROLL-Tandem Repeat Occurrence LocatorabstractSUMMARY: Tandem Repeat Occurrence Locator (TROLL), is a light-weight Simple Sequence Repeat (SSR) finder based on a slight modification of the Aho-Corasick algorithm. It is fast and only requires a standard Personal Computer (PC) to operate. We report running times of 127 s to find all SSRs of length 20 bp or more on the complete Arabdopsis genome--approx. 130 Mbases divided in five chromosomes--using a PC Athlon 650 MHz with 256 MB of RAM. AVAILABILITY: TROLL is an open source project and is available at http://finder.sourceforge.net. Adalberto T. Castelo, Wellington Santos Martins, Guang R. Gao |
Bioinform. | 2 |
| 1998 | The performance of a selection of sorting algorithms on a general purpose parallel computerabstractIn the past few years, there has been considerable interest in general purpose computational models of parallel computation to enable independent development of hardware and software. The BSP and related models represent an important step in this direction, providing a simple view of a parallel machine and permitting the design and analysis of algorithms whose performance can be predicted for real machines. In this paper we analyse the performance of three sorting algorithms on a BSP-type architecture and show qualitative agreement between experimental results from a simulator and theoretical performance equations. © 1998 John Wiley & Sons, Ltd. Roy D. Dowsing, Wellington Santos Martins |
Concurr. Pract. Exp. | 2 |